Skip to main content
llama-server is the primary inference server. It exposes an OpenAI-compatible HTTP API and an integrated web UI. Example launch command